Molecular Genetics and Genomics
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Molecular Genetics and Genomics's content profile, based on 12 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Vedder, L.; Schoof, H.
Show abstract
Biological sequences are known to be not random. Thus, the comparison of in silico restriction fragment distributions of random and biological sequences may be an indicator of this non-randomness. Our analyses show that for most of the tested combinations of restriction enzyme and genome sequence the fragments per Megabase of the biological sequence deviate at least more then 10% from the corresponding random sequence. This deviation goes into both directions, i.e. clearly increased values are as common as clearly decreased values. Although there is no species- or restriction-enzyme-specific effect, a clear impact of the GC content both of the restriction site and of the genome sequence can be seen. In contrast to the random sequences, the genome sequences show distinct peaks in their fragment length distributions, hinting to repetitive elements such as transposons.
Parida, A. S.; Kumar, A.; Tiwari, B.
Show abstract
The only autonomously active transposable elements in the human genome are Long interspersed nuclear element-1 (LINE-1) elements. These elements are known to play an important role in changing the transcriptome. LINE-1 sequences affect gene regulation during post-transcription processing, along with their established role in retrotransposition. Exonization is one mechanism where the LINE-1 integrated genome undergoes alternative splicing to produce new isoforms of transcripts. Our work mainly highlights the effect of LINE-1 associated exonization, focusing on the formation of isoforms of transcripts. Using Non-small cell lung cancer (NSCLC) as a model, we conducted a detailed transcriptome study that combines splice junction profiling with gene expression data. Our results show that LINE-1 sequences are often included as exons in host transcripts, leading to the formation of new exons and their various isoforms. The events are validated by solid splice junction evidence that proves the reliability and reproducibility. In particular, it was observed that repetitive analyses revealed certain LINE-1 exonization events that were consistent. The finding indicates that LINE-1 act as recurrent sources of splice ready sequences. Though exonizations do not necessarily affect the total expression levels of genes, our study reveals that they certainly contribute to transcript diversity. The diversity of isoforms generated potentially contributes to the effects of gene function. This study is limited to NSCLC, but it is likely that the exonizations events play a crucial role in the altering RNA diversity in cancers. Therefore the study elucidates new insights into how transposable elements modify gene structure and function during cancer development.
Bellesis, A.; Li, X.; Moore-Frederick, D.; Xu, D.; Delbridge, K.; Ma, J.; Vaccaro, G.; Edward, B. A. A.; Kellogg, M.; Creeger, Y.; Okamoto, A. S.; Kaplow, I. M.
Show abstract
Immortalized cell lines are widely used in biological research despite their known differences from their tissues and cell types of origin. Such cell lines are especially popular for testing hypotheses regarding the activity of cis-regulatory elements (CREs) that regulate gene expression. Previous investigations of blood and skin cell lines revealed many differences between the transcriptional regulatory networks of the cell lines and the associated primary cells. Similar comparisons for other tissues have been limited. Here, we used ATAC-seq to profile CREs in four immortalized liver cell lines and found many differences between each cell lines CREs and primary liver tissue, including differences in the transcription factors that are likely to bind them and differences in the genes that they are likely to regulate. Modifying cell culture conditions based on recommendations in the literature did not improve the similarity with primary liver tissue. Our results suggest that differences between the transcriptional regulatory networks in cell lines and primary tissue should be considered when designing and interpreting cell line experiments.
Mohanta, T. K.
Show abstract
Codon usage bias is a fundamental genomic characteristic that prefers non-random preferential use of synonymous codons. It is a major determinant of translational efficiency, gene regulation, and molecular evolution. However, the evolutionary bias and functional relevance of codon usage bias across the plant lineage is poorly defined and yet to understand what are the major factors responsible for relative synonymous codon usage (RSCU) in genomes and how codon usage bias influences the gene regulation, molecular evolution genomes. A genome-wide codon usage bias study of coding DNA sequences of 262 plant genome was conducted. It encompassed more than 4.6 billion codons from > 11 million coding sequences. Relative synonymous codon usage, codon adaptation index, codon-anticodon mapping, effective number of codon (ENC)-GC3, GC1,2-GC3, parity rule 2 (PR2-bias), molecular economy, and machine learning approaches were used for the study. It was found that codon usage bias was strongly non-random and exhibited a clear phylogenetic structuring. The higher plants favoured A/T-ending, whereas early-diverging lineages were enriched in G/C-ending codons. Analysis of RSCU, codon adaptation index, and codon-anticodon pairing indicated that translational selection is mediated by tRNA availability, contributing sustainability to these molecular patterns. Machine-learning approaches identified a small subset of codons having outsized influence on genome-wide codon usage landscapes. Further studies revealed the presence of robust inverse relationships between the effective number of codons and GC content at synonymous third positions. Neutrality analysis revealed approximately 61% of variation was driven by mutational pressure, tempered by selective constraints. Phylogenetic reconstruction showed a progressive relaxation of codon bias from algae to angiosperms while maintaining a conserved molecular economy cost of ~ 30 ATP per codon across the lineages. The study revealed codon usage bias is lineage-specific evolutionary conserved trait governed by mutation, selection, and translational optimization.
Mandic, K.; Hrsak, D.; Uljanic, F.; Lenhard, B.; Baresic, A.
Show abstract
Genome-wide association studies (GWAS) are the key tools for the discovery of associations between single nucleotide polymorphisms (SNPs) and phenotypic traits and have been successfully applied to many diseases and disorders. However, a great challenge is to find the gene affected by the non-coding fraction of SNPs, especially if the gene is distal in terms of genomic distance. In this study, we present a novel approach, named targPred, which utilises genomic regulatory blocks (GRBs) for inference of a connection between a certain SNP/locus and the target gene located in the same GRB, in a more robust and generalisable manner. We identified that many disease traits such as cancer and psychiatric disease have a propensity for long-range regulation. Furthermore, we showcased a childhood obesity locus which is connected to the distal BDNF gene. Finally, we propose a new web-based service based on enhancer-promoter association, to facilitate finding the causal genes for a wide array of traits and conditions.
Tantry, S. V.; Ahrendt, S.; He, G.; LaButti, K.; Lipzen, A.; Barry, K.; Culley, D.; Magnuson, J.; Spatafora, J. W.; Grigoriev, I. V.
Show abstract
The Agaricomycotina accounts for roughly a third of all described fungi. They are important due to their wide range of lifestyles and economic and environmental relevance. Certain agaricomycetes act as lignocellulose degraders, playing a significant role in forest ecosystems and bioremediation processes. These wood-decaying fungi have historically been classified as mostly white- or brown-rot based on their ability to degrade lignin, with white-rot fungi possessing a collection of lignocellulose-degrading enzymes, which are reduced or absent in brown-rot fungi. Here, we sequenced and annotated the genome of the agaricomycete Crepidotus cesatii CBS 511.95 and explored its genome and predicted enzymatic content in a comparative context. The 36.04 Mbp genome is in 235 scaffolds, with 3.34% repeat content and 12,891 predicted genes. We found that the PFAM distributions of identified orthogroups suggested that C. cesatii shows patterns more similar to white-rot fungi compared to brown-rot fungi. Additionally, C. cesatii contained multiple copies of CAZymes CBM1 and AA9 involved in hydrolysis of lignocellulose, similar to white-rot fungi. On the other hand, according to the Conserved Unique Peptide Patterns (CUPP) data for AA2 peroxidases, the key enzymes in lignin degradation, C. cesatii is more similar to brown-rot fungi. Based on our analyses we predict that C. cesatii is another representation of the continuum of wood decaying modes between white and brown rot fungi combining genetic features of both types of fungi.
Deka, N.; Beura, P. K.; Sen, P.; Aziz, R.; Kashyap, A.; Keot, D.; Jain, M.; Namsa, N. D.; Deka, R. C.; Feil, E.; Satapathy, S. S.; Ray, S. K.
Show abstract
BackgroundMutation is thought to arise mainly during replication, though transcription is also known to be mutagenic. Considering the recent reports regarding genome-wide transcription-induced mutagenesis, a distinct demonstration of specific mutation being replication-dependent and/or transcription-dependent in genomes is yet to be established. Here, we studied synonymous single-nucleotide polymorphisms (SNPs) in 2091 individual coding sequences (CDS) in the leading strand (LeS) and the lagging strand (LaS) of the Escherichia coli chromosome by comparing across 157 strains. The frequencies of complementary transitions (ti) and complementary transversions (tv) were compared in each CDS to assess parity violation in the strands. ResultsThe C[->]T and G[->]A exhibited the maximum frequency as well as the most prominent strand inequality as these tis were influenced both by the strands as well as by the expression. Interestingly, inequality between T[->]C and A[->]G was expression-dependent but strand-independent. A[->]T and G[->]T tvs were universally more frequent than their complementary T[->]A and C[->]A tvs, respectively. ConclusionsOur study demonstrates strand-independent but expression-dependent synonymous SNP inequality in CDS, supporting the role of transcription-induced mutagenesis contributing to strand inequality in the E. coli chromosome.
Pudasaini, R.; Li-Byarlay, H.; Farrell, M. C.
Show abstract
Biting behavior is an important natural defense mechanism in honeybees (Apis mellifera) against Varroa destructor. Significant variation in this behavior exists across genetic lines of honeybees, with certain colonies exhibiting higher mite-biting activity than others. Selective breeding for enhanced biting behavior provides a promising strategy for sustainable mite control and colony resilience. However, successful implementation of such breeding programs requires a comprehensive understanding of the genomic mechanism underlying this trait. In this study, RNA-seq analysis of mandible transcriptomes of 1-day and 8-day old worker honeybees from high mite biting (HB) and low mite biting (LB) colonies were performed. A total of 9,345 genes (97.30%) showed a significant differential expression between LB and HB honeybees across different ages (one-way ANOVA, FDR < 0.05). Comparison of LB vs. HB workers collected on day 1 detected 166 down-regulated and 403 up-regulated genes, whereas workers collected on day 8 identified 82 down-regulated and 46 up-regulated genes. Furthermore, the Weighted Gene Coexpression Network Analysis (WGCNA) exhibited the brown and yellow modules with significantly higher expression in HB compared with LB on both day 1 (FC = 1.52, FDR = 0.0019 and FC = 1.47, FDR = 2.7 x 10-4, respectively) and day 8 (FC = 1.27, FDR = 0.056 and FC = 1.16, FDR = 0.073, respectively). Gene Ontology (GO) enrichment analysis identified over-representation of biological processes involved in muscle contraction, chitin binding, neural signaling, oxidative phosphorylation, neuron development, electron transport chain, mitochondrial ATP synthesis, stress response, sensory perception and metabolic processes. Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways analysis identified significant enrichment of various pathways including oxidative phosphorylation, cytoskeleton-related pathways, carbon metabolism, motor proteins, ribosome-associated pathways, and citrate cycle (TCA cycle). The present findings demonstrate mite-biting behavior is associated with coordinated activation of neural, energetic and muscular system rather than a single molecular mechanism. These findings provide a basis to improve honeybee health, enhance resistance to Varroa and other ectoparasite, and eventually support sustainable beekeeping and agricultural pollination systems.
Kim, H.; Cheong, K.; Jeon, J.; Choi, G.; Koh, J.; Song, H.; Hue, Y.; Nam, Y.; Choi, B.; Lim, Y.-J.; Choi, J.; Kim, K.-T.; Lee, Y.-H.
Show abstract
Magnaporthe oryzae, the rice blast fungus, plays a role as a model organism for molecular plant-microbe interaction research. Studies on the pathogenic mechanism of this fungus revealed many genes involved in signaling pathways. As multi-omics data are being available, genomic-level researches have been conducted to uncover the underlying biological processes during the pathogenesis of M. oryzae. Identifying the genome-wide protein-protein interaction (PPI) network is one of the omics-level approaches, which helps to understand signaling and regulatory pathways. However, existing biological network resources of M. oryzae are not sufficient to decipher pathogenesis mechanisms due to the abundance of false positives/negatives. In this study, a reliable PPI network database of M. oryzae, MagNet, was constructed with three methods, including homology-based Interolog search, co-expression network construction, and domain-domain interaction (DDI)-based prediction. With three approaches altogether, the pan-network with 5,600,976 interactions was generated, including 217,531 highly confident interactions supported by all three methods. Experimental data on M. oryzae PPIs supported that our PPI network can predict PPIs with higher accuracy compared to the previously constructed databases. MagNet would provide integrated biological network data, which can help to understand the molecular mechanisms of the rice blast fungus. The PPI data can be accessed via https:/magnet.scnu.ac.kr.
Horscroft, C.; Collins, A.; Pengelly, R. J.
Show abstract
BackgroundRecombination rates can be estimated across the genome, underpinning genetic analyses such as identification of regions under selection. Accurate recombination mapping requires observing a large number of recombination events, necessitating large sample sizes to achieve high resolution. This can be prohibitive to some analyses, so population-based estimates can also be used, leveraging the increasing population genomic data available to researchers. ObjectiveThis study aimed to determine the extent to which population-based recombination maps from different human populations are similar and assess to what extent they can be used interchangeably. MethodsWavelet analysis was employed to decompose recombination rate signals along a chromosome and evaluate the proportion of variance explained at different scales. This method also enabled the assessment of correlations across scales and identification of regions with high coherence between datasets. The analysis focused on a region of chromosome 22 in human populations of European and African ancestry. ResultsRecombination rates are not closely conserved across populations, with the greatest divergence observed at fine scales. Coherence between populations varied significantly across all scales. ConclusionAs recombination maps differ substantially between human populations, for genetic analyses involving recombination maps, it is recommended to use maps specific to the population under study.
Famakinde, D. O.; Lonergan, C.; Gobert, G.; Wells, D.; McVeigh, P.
Show abstract
RNA interference (RNAi) is a widely exploited reverse-genetics tool with potential uses for disease control. Successful RNAi has been reported in trematode-vectoring snails, but the composition of RNAi effector-encoding gene complements, a key driver for RNAi efficiency, remain unstudied in these species. Using bioinformatics and comparative genomics, we searched for orthologues of 115 RNAi effector sequences in genomes or transcriptomes of four snail vectors: Biomphalaria glabrata, B. pfeifferi, Bulinus truncatus, and Lymnaea staginalis. Gene expression patterns of selected RNAi effectors were then examined across developmental stages and tissues of the model B. glabrata snail. At least 74 RNAi-related proteins were conserved across all four species, including core components known to be essential for gene silencing. Classical systemic RNAi-deficient (SID) genes that facilitate systemic RNAi in other systems were absent, suggesting that alternative pathways may compensate for dsRNA uptake and transport. Core effectors of secondary RNAi amplification and heritable RNAi were not detected. Expressions of Dicer-1, Argonaute-2, and the exonuclease Eri-1 did not vary significantly with snail size. A putative RNAi-inhibiting Staufen orthologue showed elevated expression in the ovotestis, while another putative cholesterol-interacting gene was overexpressed in the trunk tissue and may partly contribute to RNAi import. Altogether, our results present the most comprehensive overview of RNAi pathway effectors in major intermediate snail hosts for trematodes. The findings underscore the likely broad potential for RNAi use in trematode intermediate hosts as an experimental tool and potential control method.
Monteiro de Barros, M. R.; Bosch, K.; Soualhi, S.; Issa Bhaloo, S.; Chu, T.; Hemrajani, T.; Cho, J.; Ozuner, K.; Fu, R.; Geiger, H.; Robine, N.; Carter, J. E. B.; Maniatis, S.; Ryeom, S.; Tavare, S.; Nowicki-Osuch, K.
Show abstract
Background & AimsGastric epithelial cells maintain homeostasis through dynamic self-renewal mechanisms involving stem and progenitor cells; however, identifying them has been challenging. This study aims to identify stem cells of healthy gastric epithelium and cell type-specific regulators defining gastric epithelial homeostasis via single-nucleus multiome analysis. MethodsTen unique gastric samples were collected from 8-12 week old wildtype mice. Isolated nuclei were subjected to simultaneous profiling of gene expression and chromatin accessibility. After quality control, 31,598 cells were analyzed with Seurat and Signac using weighted-nearest neighbors analysis for joint RNA and ATAC clustering. Furthermore, SCENIC+, MultiVelo, EpiCHAOS and Cell plasticity score were used to uncover gene regulatory networks, cell state dynamics and lineage trajectories. ResultsOur analyses were validated by the identification of known regulators of stem-cell differentiation into mature cell types. More importantly, it revealed previously uncharacterized regulatory networks comprising novel transcription factor combinations that define cell identities, including Ppara, Pparg, Arid5b and Sox5 as candidate regulators of parietal, foveolar, chief and neck cells, respectively. Further, our data support the identity of isthmus cells as stem-like cells of healthy gastric epithelium, as evidenced by epigenetic plasticity that simultaneously contains open chromatin states of all differentiated cell types in the absence of transcriptional reprogramming. ConclusionConsistent with Waddingtons epigenetic landscape hypothesis, gastric epithelial homeostasis is controlled by orchestrated epigenetic and transcriptional programs. Contrary to the prevailing hypothesis, stem cells can be defined not by a separate epigenetic state but by epigenetic superposition of differentiated cell states. Future work is needed to define the universality of these results.
zhang, y.; Wang, D.; Zhao, R.; Li, S.; Zheng, X.; Hu, G.
Show abstract
Rana dybowskii is distributed across Northeast Asian and represents a valuable medical resource. A high-quality assembly of the genome has not yet been reproted. This species has 2n=24 chromosomes, but a huge genome size that estimated at 3.5 ~4.6 Gb in the previous studies. The relatively large chromosome size, exceeding hundreds of megabases, may result in difficulties of obtaining a complete chromosome level genome. Here, we constructed a chromosome-level genome assembly of R. dybowskii by integrating PacBio HiFi long-read sequencing for de novo assembly and CiFi (3C coupled with HiFi sequencing) for scaffolding. The final assembly consists of 12 chromosomes with a total of 3.95 Gb and a scaffold N50 length of 455 Mb. BUSCO assessment using the tetrapoda_odb12 database identified 94.2% complete and 0.5% fragmented orthologs, suggesting a high level of completeness of the assembly. Genomic annotation revealed that repetitive sequences comprise over 53% of the assembly, with retroelements and DNA transposons accounting for 22% and 25%, respectively. A total of 43,999 protein-coding genes were predicted with the assistance of RNA-seq reads from four tissues (muscle, eye, testis and skin). This high-quality chromosome-level reference genome provides a valuable genomic resource for advancing genetic studies of the species.
Hasenklever, J. C.; Paderi, V.; Hasenklever, D.; Axmann, I. M.; Schipper, K.
Show abstract
BackgroundThe corn smut fungus Ustilago maydis is an important microbial model organism representing a genetically amenable and readily cultivable basidiomycete. Research in this fungus addresses a broad range of fundamental questions and its biotechnological exploitation is on the rise. Although genetic engineering in principle is well established, efficient methodology for synthetic biology approaches such as metabolic engineering or pathway transplantation has remained limited. ResultsHere, we present a comprehensive toolbox for U. maydis based on modular cloning and the characterization of more than 20 promoters. Careful comparative evaluation of insertion loci and terminator as well as reporter effects was conducted and a novel color-based strategy for straightforward genome integration was implemented. Moreover, the cloning and subsequent one-step integration of four transcriptional units into U. maydis was demonstrated by creating a "rainbow" strain producing four fluorescent proteins. ConclusionOverall, this next generation toolkit strongly advances genetic engineering and systems biology approaches in U. maydis, fostering its development into a valuable and competitive fungal chassis and prime model, particularly in applied research.
Tao, T.; Li, P.; Zhu, Y.; Zhang, S.; Zhang, M.; Lascoux, M.; Chen, J.
Show abstract
Demographic factors are intrinsically crucial to evaluate species' extinction risk. However, measuring them remains difficult and time-consuming and the use of genomic summary statistics has been advocated to assess the conservation status of a species. In the present study, we estimated (i) the census number (Nc), (ii) effective population size (Ne) over three different time periods, recent, historical and ancient, (iii) neutral genetic diversity ({pi}4), and (iv) a measure of the efficacy of purifying selection ({pi}0/{pi}4) for 101 plant species using population genomic sequencing data. Twenty-one species are from the Plant Species with Extremely Small Populations (PSESP) program of SW China. Threatened species exhibited significantly lower Ne, Nc, {pi}4, and weaker purifying selection, but had a higher Ne/Nc ratio than non-threatened ones. Nc was the main determinant in identifying conservation status, and contemporary neutral genetic diversity was predominantly influenced by historical Ne. In the absence of demographic information, genetic parameters are a good proxy of conservation status, likely because currently threatened species also had a low historical population size. In summary, our findings suggest that direct estimates of Nc are more useful than {pi}4, although the latter remains a valuable conservation indicator. Hence, efforts such as the PSESP should be extended.
Pathmendra, P.; Enguita, F. J.; Byrne, J. A.
Show abstract
Numbers of research articles studying circRNAs have increased rapidly since 2017. Previous analyses of human circRNA articles in two high impact factor cancer research journals identified papers with wrongly identified nucleotide sequence reagents and circRNAs whose identities could not be independently verified. In the present study, verification of human nucleotide sequence reagent and cell line identities in retracted circRNA articles published from 2017-2021 in high impact factor journals found wrongly identified nucleotide sequences and/or cell lines in all 13 retracted papers. Similar analyses of human circRNA papers published in high impact factor journals in 2022 found wrongly identified, non-verifiable and/or questionable reagents in 71% (84/118) papers, where 51% (60/118) papers described at least one wrongly identified reagent. When individual error types and features of concern were considered, 2022 circRNA papers described wrongly identified nucleotide sequence reagents (52/118, 44%), questionable circRNA probes that did not meet accepted targeting requirements (34/118, 29%), non-verifiable nucleotide sequences (25/118, 21%), wrongly identified cell lines (22/118, 19%), and/or non-verifiable cell line identifiers (6/118, 5%). In summary, wrongly identified, non-verifiable and/or questionable reagents were unexpectedly frequent in human circRNA papers in high impact journals, highlighting the need for critical engagement with the circRNA literature.
Mei, C.; Ness, J.; Nakai, K.; Wunderlich, Z.
Show abstract
Developmental processes depend on carefully coordinated gene expression. Expression is modulated by the binding of transcription factors (TFs) to cis-regulatory elements (CREs), like enhancers and promoters. Many computational and experimental approaches have been developed to find CREs, particularly enhancers, in the genome, each with strengths and caveats. Given the increasing availability of ATAC-seq data and methods to find TF binding therein, we hypothesized that we could use TF footprinting tools to find clusters of TF binding events within accessible chromatin that may act as CREs. Using Drosophila anterior-posterior patterning network as a test bed, we used a digital genomic footprinting tool (DGT), TOBIAS, on previously published early embryo ATAC-seq data to characterize the TF footprint landscape of 16 TFs essential for embryonic patterning. Even in this system, with its extensive enhancer annotation, most footprinted TF binding sites lie outside of known enhancers, with intergenic and intronic regions hosting the highest TF footprint count, albeit at low density. To find potential novel enhancers, we identified high-density TF footprint clusters that are highly conserved and overlap with active enhancer histone mark signals. Five high confidence candidates were selected for reporter assay validation and all five were found to drive spatially patterned expression in the embryo. This study shows that even in a highly characterized system, the analysis of footprinted TF binding sites in ATAC-seq data can uncover new regulatory regions and suggests this approach may be helpful in using existing ATAC-seq data to find novel CREs. ARTICLE SUMMARYGiven the increasing availability of ATAC-seq datasets, workflows to exploit the data to uncover new cis-regulatory elements (CREs), including enhancers, are valuable. Using early anterior-posterior patterning in the Drosophila embryo as a test case, we find that previously published transcription factor footprinting tools and ATAC-seq data can be analyzed to yield new candidate CREs. Experimental validation confirms the activity of selected candidate CREs, suggesting that existing data can be analyzed to find novel regulatory elements.
Longoria, K. D. D.; Stroebel, B.; Gadgil, M.; Weiss, S.; Lewis, K. A.; Perez, N.; Flowers, E.
Show abstract
BackgroundWomen are disproportionately affected by multimorbid depression and type 2 diabetes (T2D), with prevalence peaking during midlife (40-64 years), a biologically dynamic timeframe due to changes associated with reproductive aging. Yet, phenotypic and mechanistic factors contributing to midlife womens disproportionate risk for co-occurrence remain poorly defined. We previously identified co-expressed microRNAs (miRs) in midlife women with prediabetes that increased odds of assignment to a high psychometabolic risk phenotype. Here, we extend these findings by characterizing putative mRNA targets of these co-expressed miRs and pathways overrepresented among mRNAs, providing insights into potential mechanisms underlying psychometabolic risk in midlife women. MethodsThis study included baseline data from midlife women (ages 40-64 years) with prediabetes who participated in the Diabetes Prevention Program (DPP) (n = 603). In silico analyses were performed using miRTarBase to identify mRNAs regulated by 3 or more of the miRs that most prominently loaded a principal component previously identified to increase odds of assignment to a high psychometabolic risk phenotype defined in this sample. Pathway enrichment analysis was conducted to assess for overrepresentation of KEGG pathways among predicted mRNA targets. To enhance interpretability, pathways were thematically clustered based on their evidenced role in human physiology. ResultsWe identified a total of 13 mRNAs targeted by co-expressed miRs associated with increased odds of assignment to a high psychometabolic risk phenotype in midlife women with prediabetes. Pathway enrichment analysis revealed a total of 71 KEGG pathways with overrepresentation of identified mRNA targets. Four overarching biological themes emerged, reflecting involvement of metabolic, inflammatory, endocrine, and stress/biological weathering-related processes. ConclusionsExperimentally validated mRNA targets related biological pathways were identified, providing multisystem insights into potential mechanisms underlying risk for multimorbid depression and T2D in midlife women. Findings offer mechanistic targets for experimental validation and future precision health research focused on this high-risk population. Overall, this work positions the utility of miRs as context-sensitive biomarkers in the characterization of risk for complex, multimorbid conditions in women during biologically dynamic timeframes.
Shvetcov, A.; Thomson, S.; Finney, C. A.
Show abstract
Human induced pluripotent stem cell (iPSC)-based disease modelling studies are widely expected to include three to five independent donor lines to control for the contribution of donor genetic background to phenotypic variance. This convention has been formalized into major guidelines, yet no power analysis has evaluated whether these sample sizes can detect, estimate, or control for donor-level genetic effects. Here, we provide that evaluation. Using Monte Carlo simulation, closed-form confidence intervals, population genetics, and empirical resampling of transcriptomic data from iPSC lines, we show that studies with three to five donors cannot reliably detect donor-level variance, cannot estimate its magnitude with useful precision, and cannot determine whether a treatment effect generalizes across genetic backgrounds. The sample sizes required to reliably detect, estimate, or control for donor-level variance exceed 20 donors and, for many phenotypes, exceed 50, well beyond what any standard disease modelling experiment can deliver. Adding two or three donor lines to a study does not meaningfully increase statistical power, narrow confidence intervals, or establish whether a treatment effect generalizes across genetic backgrounds. The inability to control for genetic background is not a limitation of individual study design but a structural property of iPSC-based modelling. We propose that the field adopt isogenic controls for variant-specific questions and orthogonal validation against clinical datasets for generalizability, rather than treating donor number as a proxy for rigour.
Shukla, M.; Bohra, D. L.; Rao, B.; Narayan, L.; Kiran, S.; Thakur, V.
Show abstract
Genomic erosion as a manifestation of small effective population size (Ne) and consanguinity subverts long-term perpetuation of threatened species by compromising their adaptive potential; however, the integration of genomics remains limited in applied conservation efforts to guide priorities. This study combines non-invasive sampling, double-digest Restriction site-associated DNA sequencing (ddRAD), and population-genomic analyses to assess genetic health in two vulture assemblages-mixed wild enclosure and captive breeding cohorts. Both the geographical locations exhibit signs of populations in distress: low genetic diversity and abundant intermediate-length runs of homozygosity (RoH), consistent with long-term reduced Ne plus recent demographic isolation. Our demographic model runs favoured ancient migration (AM) topology characterised by an ephemeral window of gene flow, taken over by a prolonged population separation period. The mutation quantification results from approximately 59,000 outgroup-polarised SNPs reveal higher additive burden and more homozygous-derived sites in BKN. However, this was later traced to low-impact and non-coding variants rather than a surge in the loss-of-function (LoF) alleles. The data support a genomic profile that carries an elevated risk from polygenic/aggregate deleterious burden in BKN despite a scarcity of high-impact mutations. By highlighting the disconnect between genetic resilience and demographic recovery, our results accentuate the need to incorporate genomics-informed inbreeding and monitoring programs, while also focusing on reducing anthropogenic mortality with genetic augmentation.